19 research outputs found

    The metaRbolomics Toolbox in Bioconductor and beyond

    Get PDF
    Metabolomics aims to measure and characterise the complex composition of metabolites in a biological system. Metabolomics studies involve sophisticated analytical techniques such as mass spectrometry and nuclear magnetic resonance spectroscopy, and generate large amounts of high-dimensional and complex experimental data. Open source processing and analysis tools are of major interest in light of innovative, open and reproducible science. The scientific community has developed a wide range of open source software, providing freely available advanced processing and analysis approaches. The programming and statistics environment R has emerged as one of the most popular environments to process and analyse Metabolomics datasets. A major benefit of such an environment is the possibility of connecting different tools into more complex workflows. Combining reusable data processing R scripts with the experimental data thus allows for open, reproducible research. This review provides an extensive overview of existing packages in R for different steps in a typical computational metabolomics workflow, including data processing, biostatistics, metabolite annotation and identification, and biochemical network and pathway analysis. Multifunctional workflows, possible user interfaces and integration into workflow management systems are also reviewed. In total, this review summarises more than two hundred metabolomics specific packages primarily available on CRAN, Bioconductor and GitHub

    Prediction, Detection, and Validation of Isotope Clusters in Mass Spectrometry Data

    No full text
    Mass spectrometry is a key analytical platform for metabolomics. The precise quantification and identification of small molecules is a prerequisite for elucidating the metabolism and the detection, validation, and evaluation of isotope clusters in LC-MS data is important for this task. Here, we present an approach for the improved detection of isotope clusters using chemical prior knowledge and the validation of detected isotope clusters depending on the substance mass using database statistics. We find remarkable improvements regarding the number of detected isotope clusters and are able to predict the correct molecular formula in the top three ranks in 92 % of the cases. We make our methodology freely available as part of the Bioconductor packages xcms version 1.50.0 and CAMERA version 1.30.0

    Combining phylogenetic footprinting with motif models incorporating intra-motif dependencies

    Get PDF
    [ Background] Transcriptional gene regulation is a fundamental process in nature, and the experimental and computational investigation of DNA binding motifs and their binding sites is a prerequisite for elucidating this process. Approaches for de-novo motif discovery can be subdivided in phylogenetic footprinting that takes into account phylogenetic dependencies in aligned sequences of more than one species and non-phylogenetic approaches based on sequences from only one species that typically take into account intra-motif dependencies. It has been shown that modeling (i) phylogenetic dependencies as well as (ii) intra-motif dependencies separately improves de-novo motif discovery, but there is no approach capable of modeling both (i) and (ii) simultaneously.[Results] Here, we present an approach for de-novo motif discovery that combines phylogenetic footprinting with motif models capable of taking into account intra-motif dependencies. We study the degree of intra-motif dependencies inferred by this approach from ChIP-seq data of 35 transcription factors. We find that significant intra-motif dependencies of orders 1 and 2 are present in all 35 datasets and that intra-motif dependencies of order 2 are typically stronger than those of order 1. We also find that the presented approach improves the classification performance of phylogenetic footprinting in all 35 datasets and that incorporating intra-motif dependencies of order 2 yields a higher classification performance than incorporating such dependencies of only order 1.[Conclusion] Combining phylogenetic footprinting with motif models incorporating intra-motif dependencies leads to an improved performance in the classification of transcription factor binding sites. This may advance our understanding of transcriptional gene regulation and its evolution.This work was financially supported by DFG (grant no. GR3526/1), Gencat (2014 SGR 118), and Collectiveware (TIN2015-66863-C2-1-R).Peer reviewe

    Unrealistic phylogenetic trees may improve phylogenetic footprinting

    Get PDF
    Motivation: The computational investigation of DNA binding motifs from binding sites is one of the classic tasks in bioinformatics and a prerequisite for understanding gene regulation as a whole. Due to the development of sequencing technologies and the increasing number of available genomes, approaches based on phylogenetic footprinting become increasingly attractive. Phylogenetic footprinting requires phylogenetic trees with attached substitution probabilities for quantifying the evolution of binding sites, but these trees and substitution probabilities are typically not known and cannot be estimated easily. Results: Here, we investigate the influence of phylogenetic trees with different substitution probabilities on the classification performance of phylogenetic footprinting using synthetic and real data. For synthetic data we find that the classification performance is highest when the substitution probability used for phylogenetic footprinting is similar to that used for data generation. For real data, however, we typically find that the classification performance of phylogenetic footprinting surprisingly increases with increasing substitution probabilities and is often highest for unrealistically high substitution probabilities close to one. This finding suggests that choosing realistic model assumptions might not always yield optimal predictions in general and that choosing unrealistically high substitution probabilities close to one might actually improve the classification performance of phylogenetic footprinting. © The Author 2017.We thank DFG [grant no. GR3526/1] for financial support.Peer Reviewe

    Prediction, Detection, and Validation of Isotope Clusters in Mass Spectrometry Data

    No full text
    Mass spectrometry is a key analytical platform for metabolomics. The precise quantification and identification of small molecules is a prerequisite for elucidating the metabolism and the detection, validation, and evaluation of isotope clusters in LC-MS data is important for this task. Here, we present an approach for the improved detection of isotope clusters using chemical prior knowledge and the validation of detected isotope clusters depending on the substance mass using database statistics. We find remarkable improvements regarding the number of detected isotope clusters and are able to predict the correct molecular formula in the top three ranks in 92 % of the cases. We make our methodology freely available as part of the Bioconductor packages xcms version 1.50.0 and CAMERA version 1.30.0

    Detecting and correcting the binding-affinity bias in ChIP-seq data using inter-species information

    Get PDF
    BACKGROUND: Transcriptional gene regulation is a fundamental process in nature, and the experimental and computational investigation of DNA binding motifs and their binding sites is a prerequisite for elucidating this process. ChIP-seq has become the major technology to uncover genomic regions containing those binding sites, but motifs predicted by traditional computational approaches using these data are distorted by a ubiquitous binding-affinity bias. Here, we present an approach for detecting and correcting this bias using inter-species information. RESULTS: We find that the binding-affinity bias caused by the ChIP-seq experiment in the reference species is stronger than the indirect binding-affinity bias in orthologous regions from phylogenetically related species. We use this difference to develop a phylogenetic footprinting model that is capable of detecting and correcting the binding-affinity bias. We find that this model improves motif prediction and that the corrected motifs are typically softer than those predicted by traditional approaches. CONCLUSIONS: These findings indicate that motifs published in databases and in the literature are artificially sharpened compared to the native motifs. These findings also indicate that our current understanding of transcriptional gene regulation might be blurred, but that it is possible to advance this understanding by taking into account inter-species information available today and even more in the future

    Untargeted In Silico Compound Classification-A Novel Metabolomics Method to Assess the Chemodiversity in Bryophytes

    No full text
    In plant ecology, biochemical analyses of bryophytes and vascular plants are often conducted on dried herbarium specimen as species typically grow in distant and inaccessible locations. Here, we present an automated in silico compound classification framework to annotate metabolites using an untargeted data independent acquisition (DIA)-LC/MS-QToF-sequential windowed acquisition of all theoretical fragment ion mass spectra (SWATH) ecometabolomics analytical method. We perform a comparative investigation of the chemical diversity at the global level and the composition of metabolite families in ten different species of bryophytes using fresh samples collected on-site and dried specimen stored in a herbarium for half a year. Shannon and Pielou's diversity indices, hierarchical clustering analysis (HCA), sparse partial least squares discriminant analysis (sPLS-DA), distance-based redundancy analysis (dbRDA), ANOVA with post-hoc Tukey honestly significant difference (HSD) test, and the Fisher's exact test were used to determine differences in the richness and composition of metabolite families, with regard to herbarium conditions, ecological characteristics, and species. We functionally annotated metabolite families to biochemical processes related to the structural integrity of membranes and cell walls (proto-lignin, glycerophospholipids, carbohydrates), chemical defense (polyphenols, steroids), reactive oxygen species (ROS) protection (alkaloids, amino acids, flavonoids), nutrition (nitrogen- and phosphate-containing glycerophospholipids), and photosynthesis. Changes in the composition of metabolite families also explained variance related to ecological functioning like physiological adaptations of bryophytes to dry environments (proteins, peptides, flavonoids, terpenes), light availability (flavonoids, terpenes, carbohydrates), temperature (flavonoids), and biotic interactions (steroids, terpenes). The results from this study allow to construct chemical traits that can be attributed to biogeochemistry, habitat conditions, environmental changes and biotic interactions. Our classification framework accelerates the complex annotation process in metabolomics and can be used to simplify biochemical patterns. We show that compound classification is a powerful tool that allows to explore relationships in both molecular biology by zooming in and in ecology by zooming out. The insights revealed by our framework allow to construct new research hypotheses and to enable detailed follow-up studies

    Effects of Water Availability in the Soil on Tropane Alkaloid Production in Cultivated Datura stramonium

    No full text
    Background: different Solanaceae and Erythroxylaceae species produce tropane alkaloids. These alkaloids are the starting material in the production of different pharmaceuticals. The commercial demand for tropane alkaloids is covered by extracting them from cultivated plants. Datura stramonium is cultivated under greenhouse conditions as a source of tropane alkaloids. Here we investigate the effect of different levels of water availability in the soil on the production of tropane alkaloids by D. stramonium. Methods: We tested four irrigation levels on the accumulation of tropane alkaloids. We analyzed the profile of tropane alkaloids using an untargeted liquid chromatography/mass spectrometry method. Results: Using a combination of informatics and manual interpretation of mass spectra, we generated several structure hypotheses for signals in D. stramonium extracts that we assign as putative tropane alkaloids. Quantitation of mass spectrometry signals for our structure hypotheses across different anatomical organs allowed us to identify patterns of tropane alkaloids associated with different levels of irrigation. Furthermore, we identified anatomic partitioning of tropane alkaloid isomers with pharmaceutical applications. Conclusions: Our results show that soil water availability is an effective method for maximizing the production of specific tropane alkaloids for industrial applications

    Chemical Diversity and Classification of Secondary Metabolites in Nine Bryophyte Species

    No full text
    The central aim in ecometabolomics and chemical ecology is to pinpoint chemical features that explain molecular functioning. The greatest challenge is the identification of compounds due to the lack of constitutive reference spectra, the large number of completely unknown compounds, and bioinformatic methods to analyze the big data. In this study we present an interdisciplinary methodological framework that extends ultra-performance liquid chromatography coupled to electrospray ionization quadrupole time-of-flight mass spectrometry (UPLC/ESI-QTOF-MS) with data-dependent acquisition (DDA-MS) and the automated in silico classification of fragment peaks into compound classes. We synthesize findings from a prior study that explored the influence of seasonal variations on the chemodiversity of secondary metabolites in nine bryophyte species. Here we reuse and extend the representative dataset with DDA-MS data. Hierarchical clustering, heatmaps, dbRDA, and ANOVA with post-hoc Tukey HSD were used to determine relationships of the study factors species, seasons, and ecological characteristics. The tested bryophytes showed species-specific metabolic responses to seasonal variations (50% vs. 5% of explained variation). Marchantia polymorpha, Plagiomnium undulatum, and Polytrichum strictum were biochemically most diverse and unique. Flavonoids and sesquiterpenoids were upregulated in all bryophytes in the growing seasons. We identified ecological functioning of compound classes indicating light protection (flavonoids), biotic and pathogen interactions (sesquiterpenoids, flavonoids), low temperature and desiccation tolerance (glycosides, sesquiterpenoids, anthocyanins, lactones), and moss growth supporting anatomic structures (few methoxyphenols and cinnamic acids as part of proto-lignin constituents). The reusable bioinformatic framework of this study can differentiate species based on automated compound classification. Our study allows detailed insights into the ecological roles of biochemical constituents of bryophytes with regard to seasonal variations. We demonstrate that compound classification can be improved with adding constitutive reference spectra to existing spectral libraries. We also show that generalization on compound classes improves our understanding of molecular ecological functioning and can be used to generate new research hypotheses
    corecore